sleep: measure the remaining time against a clock, not the kernel's remainder - #642
Conversation
|
Since you opened the PR, master has gained a monotonic clock. In 624854d, I added a new kernel import lisp_monotonic_time, and 1e88afc uses that to implement get-internal-real-time. So, get-internal-real-time is now monotonic on every platform: CLOCK_BOOTTIME on Linux, CLOCK_MONOTONIC_RAW on Darwin, QueryPerformanceCounter on Windows, and CLOCK_MONOTONIC on FreeBSD and illumos. Can you rebase and use the improved get-internal-real-time for the deadline instead of %nanosleep-clock? This would match what a similar deadline loop does in %timed-wait-on-semaphore-ptr. Some data from other platforms, from master without this PR's changes (ccl --load sleep-vs-alloc.lisp):
Darwin's nanosleep apparently works out the remaining time after the signal handler returns, so the time spent parked is counted. FreeBSD and illumos behave like Linux. |
…emainder (sleep n) returns late, and on a machine that collects often enough it does not return at all, when another thread allocates. This is issue Clozure#639. %nanosleep calls #_nanosleep and, when a signal interrupts the call, sleeps again for the remaining time the kernel wrote back. Each GC suspends the sleeping thread with a signal, and suspend_resume_handler blocks inside the handler until the collection is over. The kernel computes the remaining time before it runs the handler, so the time the thread spends blocked in the handler is not subtracted. Every world stop adds its own duration to the sleep, and the error grows with the collection rate. On two 2-vCPU burstable instances a 10 s sleep took 31.8 s on linuxx8664 and 170.1 s on linuxarm64. Earlier runs on the same arm64 cell were killed at 90 s without returning, so the overrun has no bound that I have established. Darwin does not have the defect: its nanosleep works the remaining time out after the signal handler returns, so the parked time is counted. FreeBSD and illumos behave like Linux. Those three measurements are not mine; they come from the maintainer, who ran the reproducer on platforms I cannot reach. The fix computes the deadline once and, on each EINTR, sleeps again for (stop - now). The kernel's write-back is no longer read at all, so #_nanosleep now gets a NULL rem pointer, and the Darwin check for a negative zero-extended remainder goes with the code it guarded. The clock is get-internal-real-time, which is monotonic on every platform since 1e88afc. The loop is shaped like the one in %timed-wait-on-semaphore-ptr, which solves the same problem for a semaphore wait: keep a stop in internal-time-units, re-read the clock on each wakeup, and floor the difference back into the units the call wants. clock_nanosleep with TIMER_ABSTIME was considered and not used: Darwin does not have it, so it would need a second loop with a different error convention for one target. This form is one loop for every target. Cost: two clock reads per %nanosleep call and one more per interruption. process-wait polls through %nanosleep once per tick, so this is noise there. Measured red then green with ccl-tests extended-tests/threads/sleep-vs-alloc.lisp on linuxx8664 and linuxarm64. The pull request carries the cell. Not built on a 32-bit port, and not built on Darwin, FreeBSD or illumos.
a5f01ad to
5060148
Compare
|
Rebased onto master and rewritten to use
Two things fall out of that shape. The call passes The two-timespec ping-pong goes with it. One |
(sleep n)overruns when another thread allocates, and the error grows with the collection rate. This is issue #639.%nanosleepsleeps again for the remaining time the kernel writes back on EINTR. The kernel computes that remainder before it runs the signal handler, andsuspend_resume_handlerblocks inside the handler until the collection finishes, so the time the thread spends parked is never subtracted. Every world stop adds its own duration to the sleep.There is one definition, under
#-windows-target, with no per-port override, so every non-Windows target has this. Three callers reach it:sleep, the periodic-task sleep inhousekeeping-loop, andprocess-waitonce per poll tick.The patch works the deadline out once, from a clock the Lisp reads itself, then sleeps again for the remaining time it measures on each EINTR. It never reads the kernel write-back, so
#_nanosleepnow gets a NULLrem, and the Darwin negative-remainder workaround goes with the code it guarded.That gives one loop for every target. The clock is
CLOCK_MONOTONICthrough#_clock_gettime, whichlib/time.lispalready reads the same way, andgettimeofdayon Darwin, a kernel import that needs no interface-database entry. I avoidedclock_nanosleepwithTIMER_ABSTIMEbecause Darwin does not have it, which would mean a second loop with a different error convention for one target.Measured
Red then green on both architectures at
6526e21c, changing only this patch, withextended-tests/threads/sleep-vs-alloc.lispfrom Clozure/ccl-tests#10.On arm64 the kernel md5 is identical in both halves, as it should be, since the patch touches no kernel file. Both machines are 2-vCPU burstable instances, so read those as verdicts and not as benchmark figures.
ANSI reports 21679/0 and ccl-tests 243/0 with the patch applied.
I have not tested Darwin. There is no Darwin build here, so the
gettimeofdayarm rests on construction rather than measurement.